Papers with Automatic Speech Recognition

90 papers
Optimizing Entity Resolution in Voice Interfaces: An ASR-Aware Entity Reference Expansion Approach (2024.emnlp-industry)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) errors in voice-based dialog systems pose significant impediments to downstream tasks.
Approach: They propose an automatic speech recognition (ASR) error-aware loss function to inject failed mentions and resolved entity names into the knowledge graph to enhance its awareness of unresolved mentions.
Outcome: The proposed system enhances the knowledge graph's awareness of unresolved mentions by injecting pairs of failed mentions and resolved entities into the knowledge map.
Advancing African-Accented English Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models (2025.acl-srw)

Copied to clipboard

Challenge: Accents play a pivotal role in shaping human communication, a new study finds . existing ASR systems often perform inadequately, even mispronouncing African names .
Approach: They propose a method that uses epistemic uncertainty to automate annotation to reduce costs and human labor.
Outcome: The proposed method reduces costs and human labor by reducing data annotation and epistemic uncertainty.
metaCAT: A Metadata-based Task-oriented Chatbot Annotation Tool (2020.aacl-demo)

Copied to clipboard

Challenge: Creating high-quality annotated dialogue corpora necessitates a high level of human engagements.
Approach: They propose to develop an annotation tool specifically for developing task-oriented dialogue data that provides comprehensive metadata annotation coverage to the domain, intent, and span information.
Outcome: The tool provides comprehensive metadata annotation coverage to domain, intent, and span information.
Hyper-BTS Dataset: Scalability and Enhanced Analysis of Back TranScription (BTS) for ASR Post-Processing (2024.findings-eacl)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) post-processing requires substantial amounts of data, requiring expensive phonetic transcription experts.
Approach: They propose a "Hyper-BTS" dataset that is five times larger than prior studies . they propose criteria for categorizing error types within ASR post-processing .
Outcome: The proposed method can generate ASR inputs from clean text using a text-to-speech system.
Audio Query Handling System with Integrated Expert Models and Contextual Understanding (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing chatbots are limited to specific audio tasks, but the domain of audio content related queries remains underexplored.
Approach: They propose to use an intent classifier to route queries to audio-related experts using a diverse audio query dataset.
Outcome: The proposed system outperforms state-of-the-art LLMs on custom audio tasks and MMAU sound set benchmarks.
Growing Trees on Sounds: Assessing Strategies for End-to-End Dependency Parsing of Speech (2024.acl-short)

Copied to clipboard

Challenge: Direct dependency parsing of the speech signal is proposed as a way of incorporating prosodic information into the parser and bypassing the limitations of a pipeline approach.
Approach: They propose to use graph-based parsing and sequence labeling based parses to integrate prosodic information into the parser and bypass limitations of pipeline approaches.
Outcome: The proposed graph based approach outperforms a pipeline approach on a large treebank of spoken french, despite having 30% fewer parameters.
No Label? No Problem: Unsupervised Continual Learning for Adaptive Medical ASR (2026.eacl-industry)

Copied to clipboard

Challenge: Medical audio often contains specialized terminology, such as medication names, which existing ASR systems struggle to transcribe accurately.
Approach: They propose an unsupervised continual learning ASR framework that adapts to new data while preserving prior knowledge.
Outcome: Experiments on real-world medical audio show that the proposed framework improves over state-of-the-art models.
Audio De-identification - a New Entity Recognition Task (N19-2)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is an important step in de-identification (de-ID) of medical records, many of which are recorded conversations between a patient and a doctor.
Approach: They propose to use Named Entity Recognition (NER) to detect audio spans with entity mentions in medical records and then use it to evaluate the results.
Outcome: The proposed pipeline is based on a large labeled segment of the Switchboard and Fisher audio datasets and compares it with a benchmark.
Building Accurate Low Latency ASR for Streaming Voice Search in E-commerce (2023.acl-industry)

Copied to clipboard

Challenge: Recent years have witnessed the popularity of end-to-end ASR models, which have demonstrated higher accuracy compared to traditional pipelines with separate acoustic, pronunciation, and language models.
Approach: They build accurate LSTM, attention and CTC based streaming ASR models for large-scale Hinglish voice search.
Outcome: The proposed model achieves a word error rate (WER) of 3.69% without EOS and 4.78% with EOS, with 1300 ms (46.64%) reduction in latency.
Exploring the Effect of Dialect Mismatched Language Models in Telugu Automatic Speech Recognition (2022.naacl-srw)

Copied to clipboard

Challenge: Existing studies have found that the ASR system is susceptible to dialect variations within a language, thereby adversely affecting the APR.
Approach: They propose to build a dialect-specific AM while keeping the Language Model constant for all the dialects and to reduce the degradation by 9% and 15%.
Outcome: The proposed model can be built for three different Telugu regional dialects while keeping the Language Model constant for all the dialects.
A Recorded Debating Dataset (L18-1)

Copied to clipboard

Challenge: Existing research in computational argumentation and debating technologies focuses on argumentation mining, but other tasks are being addressed as well.
Approach: They describe a dataset of debating speeches in English that is used for research . they use an automatic speech recognition system to produce a more "nLP-friendly" text .
Outcome: The proposed dataset contains 60 speeches on various controversial topics, each in five formats corresponding to different stages in production.
Synthetic Doctor-Patient Dialogue Generation for Robust Medical ASR: A Scalable Pipeline for Vocabulary Expansion and Privacy Preservation (2026.eacl-industry)

Copied to clipboard

Challenge: Existing ASR models struggle with high word error rates (WER) on clinical vocabulary, especially medication names.
Approach: They propose to generate doctor-patient dialogues in both text and audio formats using a curated set of over 124,000 medical terms.
Outcome: The proposed pipeline generated over 1 billion audios with ground truth transcriptions.
Post-ASR Correction in Hindi: Comparing Language Models and Large Language Models in Low-Resource Scenarios (2026.eacl-short)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems for low-resource languages produce erroneous transcripts due to limited annotated data and linguistic complexity.
Approach: They compare language models and large language models for post-ASR correction in Hindi . they observe a scaling trend under zero-shot ICL where mid-sized LLMs degrade performance before marginal recovery at extreme scales.
Outcome: The proposed model outperforms larger models in both fine-tuning and in-context learning settings.
PriMock57: A Dataset Of Primary Care Mock Consultations (2022.acl-short)

Copied to clipboard

Challenge: Recent advances in Automatic Speech Recognition (ASR) have made it possible to reliably produce automatic transcripts of clinician-patient conversations.
Approach: They present a public access, high quality dataset of 57 mocked primary care consultations . they aim to offer a benchmark for conversational medical ASR and consultation note generation from transcripts.
Outcome: The proposed dataset can be used as a benchmark for conversational medical ASR and consultation note generation from transcripts.
Thesis Proposal: Self-Adaptive and Epistemic Uncertainty-Guided ASR of Dense Intra-Sentential Code-Switched Speech for African Low-Resource Languages (2026.acl-srw)

Copied to clipboard

Challenge: Existing multilingual and pretrained ASR systems improve general recognition accuracy but are weak at switch regions and are sensitive to language imbalance during adaptation.
Approach: They propose a self-adaptive and epistemic uncertainty-guided framework for African low-resource code-switched ASR using Hausa–English and Hausa-Yorùbá as case studies.
Outcome: The proposed framework is based on Hausa–English and Hausa-Yorùbá as case studies.
Who Needs Decoders? Efficient Estimation of Sequence-Level Attributes with Proxies (2024.eacl-long)

Copied to clipboard

Challenge: Autoregressive decoding is expensive for many sequence-to-sequence tasks, but for some downstream tasks, the actual decoding output is not needed, just attributes of the sequence.
Approach: They propose non-autoregressive proxy models that can efficiently predict scalar-valued sequence-level attributes from the encodings, avoiding the expensive decoding stage.
Outcome: The proposed models outperform ensembles in machine translation (MT) and automatic speech recognition (ASR) while being significantly faster.
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction (2025.emnlp-industry)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems have a substantial number of erroneous recognition due to environmental noise, ambiguity, etc.
Approach: They use a benchmark dataset to analyze ASR errors in the Chinese language . they then apply large language models to correct ASR error correction .
Outcome: The proposed method is based on a dataset of ASR errors in the Chinese language . it shows prompting is not effective for ASR error correction .
A Dataset for Speech Emotion Recognition in Greek Theatrical Plays (2022.lrec-1)

Copied to clipboard

Challenge: Speech Emotion Recognition (SER) is a task that is difficult to perform by humans due to subjectiveness of the emotional content.
Approach: They propose to use GreThE to collect data for speech emotion recognition in Greek plays.
Outcome: The proposed dataset contains utterances from various actors and plays, along with respective valence and arousal annotations.
Chinese Spoken Named Entity Recognition in Real-world Scenarios: Dataset and Approaches (2024.findings-acl)

Copied to clipboard

Challenge: Current Chinese Spoken NER datasets are laboratory-controlled and are limited in topics.
Approach: They propose to use Chinese Spoken NER datasets to extract entities from speech to help voice assistants better grasp the intent behind user's questions and instructions.
Outcome: The proposed methods improve on self-training-asr and mapping then distilling, and even compared with GPT4.0, they achieve better results.
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation (2025.findings-acl)

Copied to clipboard

Challenge: Generative Error Correction (GEC) is a powerful post-processing method to boost the performance of Automatic Speech Recognition systems.
Approach: They propose a method to augment GEC models with retrieved entities to improve accuracy in out-of-domain and out-od scenarios.
Outcome: The proposed method outperforms baseline models on multiple datasets and settings.
Automatic Speech Recognition System-Independent Word Error Rate Estimation (2024.lrec-main)

Copied to clipboard

Challenge: Word error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition systems.
Approach: They propose a hypothesis generation method for ASR system-dependent WER estimation . they use phonetically similar or linguistically more likely alternative words to generate hypotheses .
Outcome: The proposed method outperforms baseline estimators on in-domain data and out-of-domain on Switchboard and CALLHOME.
RED-ACE: Robust Error Detection for ASR using Confidence Embeddings (2022.emnlp-main)

Copied to clipboard

Challenge: ASR Error Detection (AED) models post-process the output of Automatic Speech Recognition systems, in order to detect transcription errors.
Approach: They propose to use ASR model's word-level confidence scores to combine ASR models with transcribed text to improve AED performance.
Outcome: The proposed models combine the confidence scores and transcribed text into a contextualized representation.
Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders (2021.acl-long)

Copied to clipboard

Challenge: End-to-end Speech Translation (E2E ST) encoders lack global context representation, whereas MT encoder lacks it.
Approach: They propose a Stacked Acoustic-and-Textual Encoding method for speech translation . they propose an adaptor module to alleviate representation inconsistency .
Outcome: The proposed method achieves state-of-the-art BLEU scores of 18.3 and 25.2 on two ST tasks.
Direct Segmentation Models for Streaming Speech Translation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to stream ST combine advances in ASR and MT to achieve high quality translations without compromising the speed of the system.
Approach: They propose to concatenate an Automatic Speech Recognition system followed by a Machine Translation system.
Outcome: The proposed models improve on the Europarl-ST dataset on the BLEU score.
BEA-Base: A Benchmark for ASR of Spontaneous Hungarian (2022.lrec-1)

Copied to clipboard

Challenge: Hungarian is spoken by 15 million people, yet, easily accessible Automatic Speech Recognition (ASR) benchmark datasets are practically unavailable.
Approach: They propose to use a subset of the BEA spoken Hungarian database to assess ASR, primarily for conversational AI applications.
Outcome: The proposed framework achieves 45% reduction in recognition error rate compared to classical approach without external language model or additional supervised data.
Isometric Neural Machine Translation using Phoneme Count Ratio Reward-based Reinforcement Learning (2024.findings-naacl)

Copied to clipboard

Challenge: Traditional Automatic Video Dubbing (AVD) pipelines use isometric-NMT algorithms to regulate the length of the output text.
Approach: They propose an isometric-NMT system that regulates the length of the output text . they propose a phoneme Count Compliance score to measure length compliance .
Outcome: The proposed approach improves phoneme count compliance scores by 36% compared to state-of-the-art models in English-Hindi language pairs.
KEBAP: Korean Error Explainable Benchmark Dataset for ASR and Post-processing (2023.emnlp-main)

Copied to clipboard

Challenge: Conventional evaluation metrics for automatic speech recognition systems produce a singular aggregate score, which is insufficient for understanding specific system vulnerabilities.
Approach: They propose to introduce the Korean Error Explainable Benchmark Dataset for ASR and Post-processing (KEBAP) this method enables a more balanced assessment encompassing speech recognition accuracy and user readability.
Outcome: The proposed method enables a more balanced assessment encompassing speech recognition accuracy and user readability.
WER we are and WER we think we are (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent reports of very low word error rates (WERs) achieved by modern automatic speech recognition systems are skepticism towards the accuracy of modern systems.
Approach: They propose to use a dataset to test automatic speech recognition systems . they propose guidelines for creating real-life datasets with high quality annotations .
Outcome: The proposed system achieves 81% of accuracy on human-chatbot interactions compared to the best reported results on human conversations and public benchmarks.
A Comprehensive Evaluation of Incremental Speech Recognition and Diarization for Conversational AI (2020.coling-main)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems are increasingly powerful and more numerous with several options existing as a service.
Approach: They evaluate the most popular automatic speech recognition systems with metrics and experiments designed with these standards in mind.
Outcome: The most popular ASR systems are Microsoft and IBM, and none are suitable for natural spontaneous conversations in real-time.
Contrastive and Consistency Learning for Neural Noisy-Channel Model in Spoken Language Understanding (2024.naacl-long)

Copied to clipboard

Challenge: End-to-end learning models require large volume of speech data with intent labels . however, models are sensitive to inconsistencies between training and evaluation conditions .
Approach: They propose a module-based approach to learn intent in a noisy-channel model . they correlate error patterns between clean and noisy ASR transcripts .
Outcome: The proposed method outperforms existing methods and improves in noisy environments.
DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation (2023.acl-long)

Copied to clipboard

Challenge: State-of-the-art automatic speech recognition systems exhibit disparate performance on varying speech accents.
Approach: They propose to use submodular mutual information to find the most informative set of utterances matching a target accent within a fixed budget.
Outcome: The proposed model is 3-5 times more label-efficient on the Indic-TTS and L2 datasets than other methods.
WER-BERT: Automatic WER Estimation with BERT in a Balanced Ordinal Classification Paradigm (2021.eacl-main)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems are evaluated using Word Error Rate (WER) a higher WER means a lower percentage of errors between the ground truth and the transcription of the system.
Approach: They propose a new balanced paradigm for automatic Word Error Rate estimation using a Librispeech dataset and a Google Cloud's Speech-to-Text API.
Outcome: The proposed approach is more effective than regression in a classification setting, but suffers from heavy class imbalance.
A Benchmark of French ASR Systems Based on Error Severity (2025.coling-main)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) transcription errors are often assessed using metrics that compare them with a reference transcription.
Approach: They propose to categorize transcription errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis.
Outcome: The proposed evaluation categorizes errors into four levels of severity based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis.
Vietnamese Automatic Speech Recognition: A Revisit (2026.findings-eacl)

Copied to clipboard

Challenge: Existing datasets with low quality and inconsistent annotations are insufficient for high-quality models.
Approach: They propose a pipeline for aggregating and preprocessing high-quality ASR datasets from diverse, potentially noisy, open-source sources.
Outcome: The proposed pipeline provides a foundation for training and evaluating state-of-the-art Vietnamese ASR systems.
Generative Music Models’ Alignment with Professional and Amateur Users’ Expectations (2025.findings-acl)

Copied to clipboard

Challenge: Recent years have witnessed rapid advances in text-to-music generation using large language models.
Approach: They propose a task to align AI-generated music with human expressions . they use a dataset of over 1.5 million songs to analyze their content .
Outcome: The proposed framework outperforms baseline models and facilitates end-to-end generation of songs audio.
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning (2025.acl-long)

Copied to clipboard

Challenge: Despite their growing importance, the quality of these datasets remains under-researched.
Approach: They propose guidelines and recommendations to address quality issues in future dataset development . they find that macro-level issues are more prevalent in less institutionalized, often under-resourced languages .
Outcome: The results highlight the need for proactive language planning and enhanced data quality control in the process of automatic speech recognition dataset creation.
Comprehensive Punctuation Restoration for English and Polish (2021.findings-emnlp)

Copied to clipboard

Challenge: Punctuation restoration is a fundamental requirement for the readability of text derived from Automatic Speech Recognition systems.
Approach: They evaluate several methods in the comprehensive punctuation reconstruction task by comparing two languages with a model to determine the quality of the punctuated word.
Outcome: The proposed model improves on two languages with relatively simple and complex morphologies.
Beyond Common Words: Enhancing ASR Cross-Lingual Proper Noun Recognition Using Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: In this work, we address the challenge of cross-lingual proper noun recognition in automatic speech recognition systems where proper nodes in an utterance may originate from a language different from the language in which the ASR system is trained.
Approach: They propose a dictionary-based method to correct ASR predictions in a large language model .
Outcome: The proposed method significantly reduces word error rates across cross-lingual proper noun recognition tasks involving three secondary languages.
DECM: Evaluating Bilingual ASR Performance on a Code-switching/mixing Benchmark (2024.lrec-main)

Copied to clipboard

Challenge: Code-switched (CSW) speech is a linguistic phenomenon that occurs when spoken utterances switch languages between sentences.
Approach: They propose to use a dataset to evaluate German-English CSW speech . they show that the dataset includes splits with varying degrees of CSW .
Outcome: The proposed dataset includes spontaneous speech from diverse domains, enabling realistic CSW evaluation in German-English.
Advancing Test-Time Adaptation in Wild Acoustic Test Settings (2024.emnlp-main)

Copied to clipboard

Challenge: Existing wild vision TTA methods fail to handle speech data due to the unique characteristics of high-entropy speech frames, which are unreliably filtered out even when containing crucial semantic content.
Approach: They propose a method for acoustic foundation models to perform confidence-based adaptation in wild acustic test settings.
Outcome: The proposed method outperforms baselines under Gaussian noise, environmental sounds, accent variations, and sung speech in the wild.
Samrómur: Crowd-sourcing Data Collection for Icelandic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Samrómur is a web application built upon Mozilla Foundation’s open-source voice collection platform Common Voice.
Approach: They describe an ongoing project of speech data collection using the web application Samrómur which is built upon Common Voice, Mozilla Foundation’s web platform for open-source voice collection.
Outcome: The proposed system will be the largest open speech corpus for Icelandic collected from the public domain.
Non-Autoregressive Chinese ASR Error Correction with Phonological Training (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods to correct ASR errors focus on fixed-length corrections, but rarely consider variable-length ones.
Approach: They propose a non-autoregressive method to correct Chinese ASR errors . they use phonological tokens to extend the source sentence for variable-length correction .
Outcome: The proposed method improves word error rate and speeds up inference by 6.2 times compared with the autoregressive model.
Discovering Canonical Indian English Accents: A Crowdsourcing-based Approach (L18-1)

Copied to clipboard

Challenge: Automated Speech Recognition systems degrade in performance when recognizing accents that are different from the ones in training data.
Approach: They propose to adapt Acoustic Models that are trained on one accent to a target accent by using a small amount of speech data in the target accent.
Outcome: The proposed model can be used to identify accents in Indian English and other languages.
Multilingual Models for ASR in Chibchan Languages (2024.naacl-long)

Copied to clipboard

Challenge: Existing algorithms for low resource-intensive languages are not available for these languages . a paper comparing the performance of different models and algorithms for these extremely low resource languages is presented.
Approach: They propose to fine-tune four ASR algorithms to create monolingual models for Bribri and Cabécar . they then use the best performing algorithm to train joint and transfer learning models for both languages .
Outcome: The proposed algorithms are effective in both Bribri and Cabécar, but especially in Bribri.
Leveraging Large Pre-trained Multilingual Models for High-Quality Speech-to-Text Translation on Industry Scenarios (2025.coling-main)

Copied to clipboard

Challenge: Speech-to-Text Translation systems rely on a sequential pipeline that combines ASR and MT models.
Approach: They propose a parameter-efficient framework that integrates one LPSM with a multilingual MT engine.
Outcome: The proposed framework integrates one LPSM with a multilingual MT engine.
Large Vocabulary Read Speech Corpora for Four Ethiopian Languages: Amharic, Tigrigna, Oromo and Wolaytta (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) is one of the most important technologies to support spoken communication in modern life.
Approach: They have developed four large speech corpora for four Ethiopian languages . they have word error rates of 37.65%, 31.03%, 38.02%, 33.89% for each language .
Outcome: The proposed corpora achieve word error rates of 37.65%, 31.03%, 38.02%, 33.89% for Amharic, Tigrigna, Oromo and Wolaytta.
Sequential Randomized Smoothing for Adversarially Robust Speech Recognition (2021.emnlp-main)

Copied to clipboard

Challenge: Existing, naive defenses against adversarial attacks are lagging . a new paper aims to break these defenses with adaptive noise ensembling .
Approach: They propose a randomized smoothing paradigm that can be used to break adversarial attacks . they use speech enhancement methods and a novel use for ASR output ensembling methods .
Outcome: The proposed model is robust to all attacks that use inaudible noise and can only be broken with very high distortion.
ArzEn: A Speech Corpus for Code-switched Egyptian Arabic-English (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of Arabic-English code-switching (CS) spontaneous speech is collected in an Egyptian university soundproof room . the language in Egypt is rather complex and poses many challenges to natural language processing (NLP)
Approach: They present an Egyptian Arabic-English code-switching (CS) spontaneous speech corpus.
Outcome: The proposed corpus is designed to be used in automatic speech recognition systems . it provides a useful resource for analyzing the CS phenomenon from linguistic, sociological, and psychological perspectives.
Semantic Role Labeling from Chinese Speech via End-to-End Learning (2024.findings-acl)

Copied to clipboard

Challenge: Semantic role labeling (SRL) has traditionally focused on text input.
Approach: They propose an end-to-end approach for SRL from speech integrating ASR and SRL in a joint-learning framework, focusing on the Chinese language.
Outcome: The proposed model improves on the Chinese Proposition Bank 1.0 dataset and the existing model with improved performance.
MCLF: A Multi-grained Contrastive Learning Framework for ASR-robust Spoken Language Understanding (2023.findings-emnlp)

Copied to clipboard

Challenge: Trending ASR-robust SLU systems have seen impressive improvements through global contrastive learning, but they can easily lead to severe semantic changes.
Approach: They propose a two-stage multi-grained contrastive learning framework to improve ASR robustness . they first adapt pre-trained language models to downstream SLU datasets and then fine-tune it on the corresponding dataset.
Outcome: The proposed framework improves on four datasets and four BERT-like backbone models.
Residual Adapters for Parameter-Efficient ASR Adaptation to Atypical and Accented Speech (2021.emnlp-main)

Copied to clipboard

Challenge: Automatic Speech Recognition systems perform poorly on atypical speech and heavily accented speech.
Approach: They add a residual adapter to the encoder layer to improve model adaptation . they show that the residual adapters update only a tiny fraction of the model parameters .
Outcome: The proposed model fine-tuning improves performance on atypical and accented speech . the system can update only a tiny fraction of the model parameters .
A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models.
Approach: They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages.
Outcome: The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data.
UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling Correction (2023.acl-long)

Copied to clipboard

Challenge: Chinese Spelling Correction (CSC) is a task of detecting and correcting misspelled charac- ters in Chinese texts.
Approach: They propose a model to learn detection and correction parts together from a multi-task learning perspective.
Outcome: The proposed model can learn detection and correction parts together from a multi-task learning perspective.
Evaluating the Efficacy of Large Acoustic Model for Documenting Non-Orthographic Tribal Languages in India (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained Large Acoustic Models have been shown to improve performance in spoken languages . however, their potential for novel under-resourced languages is not fully known .
Approach: They propose to use pre-trained Large Acoustic Models to document under-resourced languages . they use scripts from languages that hold a prominent presence in the geographical regions .
Outcome: The proposed model can document under-resourced languages in the electronic domain . the model can be used to document languages with a written script .
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)

Copied to clipboard

Challenge: Dialectal Arabic datasets embody a range of domain, dialect, and quality.
Approach: They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects.
Outcome: The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing.
Evaluating Workflows for Creating Orthographic Transcripts for Oral Corpora by Transcribing from Scratch or Correcting ASR-Output (2024.lrec-main)

Copied to clipboard

Challenge: Automated speech recognition systems can reduce transcription effort, but few studies have evaluated this potential.
Approach: They compare efforts for manual transcription vs. automatic correction of ASR-output . they use audio recordings from varying settings to create orthographic transcripts .
Outcome: The proposed methods reduce transcription time by 7 times on average for selected data and transcription conventions compared with corrected transcripts . the more complex the primary data, the more time has to be spent on corrections - the paper concludes a similar study could be conducted in 2022 .
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models (2025.acl-long)

Copied to clipboard

Challenge: Existing RAG frameworks rely on Automatic Speech Recognition to process speech input, which discards crucial audio information and increases computational overhead.
Approach: They propose a retrieval augmented generation framework with native, end-to-end audio support that integrates audio and text into a unified knowledge representation.
Outcome: The proposed framework can perform 10x faster than current pipelines while delivering 10x acceleration.
Extracting Biomedical Entities from Noisy Audio Transcripts (2024.lrec-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is particularly affected by noise, often termed the ASR-NLP gap.
Approach: They propose a dataset to bridge the ASR-NLP gap in the biomedical domain by extracting adverse drug reactions and mentions of entities from the Brief Test of Adult Cognition by Telephone (BTACT) exam.
Outcome: The proposed method can clean 2,000 clean and noisy recordings and eliminate errors using zero-shot and few-shot methods.
INTapt: Information-Theoretic Adversarial Prompt Tuning for Enhanced Non-Native Speech Recognition (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to improve ASR performance with pre-trained models require updating the pre-training model weights.
Approach: They propose a method that uses prompts concatenated to the original input to re-modulate attention of the pre-trained model.
Outcome: The proposed model improves the performance of L2 English and increases similarity between L2 and L1 accents.
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Automatic Speech Recognition (ASR) have been fueled by massive speech corpora, but extending coverage to diverse languages with limited resources remains a formidable challenge.
Approach: They propose a pipeline that converts large-scale text corpora into synthetic speech using off-the-shelf text-to-speech (TTS) models.
Outcome: The proposed pipeline generates 500,000 hours of synthetic speech in ten languages and achieves transcription error reductions of over 30%.
Effectively pretraining a speech translation decoder with Machine Translation data (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to improve the performance of AST systems are based on pretraining the encoder parameters using an ASR model, but using a pretrained MT decoder is not beneficial or improves the results.
Approach: They propose to use an adversarial regularizer to bring the encoder representations of the ASR and NMT tasks closer even though they are in different modalities.
Outcome: The proposed model can be pre-trained using the Automatic Speech Recognition (ASR) task even in different languages and improves in low resource settings.
AccentDB: A Database of Non-Native English Accents to Assist Neural Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: aaron e. sanchez and joe saunders: automatic speech recognition still faces a major challenge . they say accents are a way of pronouncing a language, and speakers always have manner of speaking . esassen: accents can be used to identify non-native speakers of a speech .
Approach: They propose to create a database of speech samples in non-native accents for ASR testing . they also propose to introduce accent neutralization of non- native accents to native accent .
Outcome: The proposed model is compared against human-labelled accent classes and is generalized against human data.
Multi-Sentence Resampling: A Simple Approach to Alleviate Dataset Length Bias and Beam-Search Degradation (2021.emnlp-main)

Copied to clipboard

Challenge: Neural Machine Translation suffers from a beam-search problem after a certain point, especially for long sentences.
Approach: They propose a data augmentation technique that concatenates several sentences from the original dataset to make a long training example.
Outcome: The proposed technique significantly reduces degradation with growing beam size and improves translation quality on the IWSTL15 En-Vi, IWStl17 En-Fr, and WMT14 En-De datasets.
To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation (2024.acl-long)

Copied to clipboard

Challenge: Existing models for multilingual automatic speech recognition (ASR) are computationallyintensive and lack proper comprehensive evaluations.
Approach: They propose to distill knowledge from large teacher models into smaller student variants that are more efficient.
Outcome: The proposed model outperforms existing models on standard benchmarks and dialectal data.
Masked Audio Text Encoders are Effective Multi-Modal Rescorers (2023.findings-acl)

Copied to clipboard

Challenge: Masked Language Models (MLMs) have proven to be effective for second-pass rescoring in Automatic Speech Recognition systems.
Approach: They propose a multi-modal masked language model rescorer which integrates acoustic representations into the input space of MLM.
Outcome: The proposed model reduces word error rate (WER) by 4%-16% on in-domain and 3%-7% on out-of-domain datasets over the text-only baseline.
Resilience of Large Language Models for Noisy Instructions (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are powerful tools for interpreting human commands and generating text.
Approach: They examine the resilience of large language models against five common types of disruptions including ASR, OCR, grammatical errors, typographical errors and distractive content.
Outcome: The models show resistance to noise, but their performance suffers . authors evaluated the models against five common types of disruptions based on their results .
PersonaLM: Language Model Personalization via Domain-distributed Span Aggregated K-Nearest N-gram Retrieval Augmentation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing language modeling tools for automatic speech recognition (ASR) are difficult to personalize.
Approach: They propose a domain-distributed Span-Aggregated K-nearest N-gram retrieval augmentation to improve language modeling for automatic speech recognition (ASR) personalization.
Outcome: The proposed model outperforms baselines on Wikitext-103, UserLibri, and ASAP datasets with a 10-16% improvement in perplexity and a 5-8% reduction in word error rates.
MASRI-HEADSET: A Maltese Corpus for Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Maltese is the national language of Malta and is spoken by approximately 500,000 people.
Approach: They present the first spoken Maltese corpus designed purposely for Automatic Speech Recognition (ASR) it consists of 8 hours of speech paired with text, recorded by using short text snippets in a laboratory environment.
Outcome: The MASRI-HEADSET corpus was developed by the MASR project at the University of Malta.
DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies have demonstrated that postprocessing speech recognition transcriptions with large language models can significantly enhance the accuracy of Automatic Speech Recognition (ASR).
Approach: They propose a method to improve Named Entity (NE) correction in Automatic Speech Recognition systems by leveraging phonetic similarity and augmented definitions.
Outcome: The proposed method outperforms baseline methods on common voice and STOP datasets and achieves a 28% reduction in WER and NE hit ratio.
Corpus Generation for Voice Command in Smart Home and the Effect of Speech Synthesis on End-to-End SLU (2020.lrec-1)

Copied to clipboard

Challenge: Massive amounts of annotated data are often unavailable for novel tasks performed in real-world environments such as smart homes.
Approach: They propose to use a synthetic semantically-annotated corpus of French commands for smart-home to train pipeline and end-to-end (E2E) SLU models.
Outcome: The proposed model trains pipeline and end-to-end (E2E) SLU models on voice commands acquired in a real smart home.
Using Automatic Speech Recognition in Spoken Corpus Curation (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) is a new way to make audio-visual data accessible.
Approach: They propose to use automatic speech recognition (ASR) to make audio-visual data accessible by systematic queries.
Outcome: The proposed system has higher recognition scores for the north of Germany vs. lower scores for south of the country.
Automatic Speech Recognition for Uyghur through Multilingual Acoustic Modeling (2020.lrec-1)

Copied to clipboard

Challenge: Low-resource languages suffer from lower performance of Automatic Speech Recognition (ASR) due to the lack of data.
Approach: They propose to use Turkish as donor language to train acoustic models using multilingual training to achieve more context coverage.
Outcome: The proposed system performs better with multilingual training for the under-resourced Uyghur language.
CEASR: A Corpus for Evaluating Automatic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems are increasingly needed for research and practical applications.
Approach: They propose to use public speech corpora to evaluate the quality of automatic speech recognition (ASR) they calculate an average Word Error Rate (WER) per corpus, per system and per corpor-system pair .
Outcome: The proposed corpus evaluates the quality of automatic speech recognition systems using public speech corpora and transcripts generated by state-of-the-art systems.
MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible (2020.lrec-1)

Copied to clipboard

Challenge: The Bible is the same for all the languages, thus constituting a multilingual and comparable 2 spoken corpus, is not exploited to date.
Approach: They propose to add multilingual links between small speech segments in different languages . they use a large dataset of 8,130 parallel spoken utterances across 8 languages - maSS .
Outcome: The proposed model can build automatic speech recognition models for 700 languages.
ÌròyìnSpeech: A Multi-purpose Yorùbá Speech Corpus (2024.lrec-main)

Copied to clipboard

Challenge: rynSpeech corpus is a dataset that can be used for both Text-to-Speecher (TTS) and Automatic Speech Recognition (ASR) speakers of many African languages have no access to voice-enabled applications in their native languages.
Approach: They propose a dataset to collect Yorùbá speech data that can be used for both TTS and ASR tasks.
Outcome: The proposed dataset can generate a good quality model with as little as 5 hours of speech . the results are consistent with previous studies on the Yorùbá language .
On Construction of the ASR-oriented Indian English Pronunciation Dictionary (2020.lrec-1)

Copied to clipboard

Challenge: Indian English (IE) has distinctive characteristics, especially phonologically, from other varieties of English.
Approach: They build a small IE spontaneous speech corpus and use a linguistically-guided IE pronunciation dictionary to apply it to IE.
Outcome: The proposed system performs better on IE spontaneous speech data than the one trained with CMUdict.
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding (2024.findings-emnlp)

Copied to clipboard

Challenge: End-to-end models for Spoken Language Understanding have been autoregressive, resulting in higher latencies.
Approach: They propose a method that uses Connectionist Temporal Classification to train robust non-autoregressive deliberation models.
Outcome: The proposed method achieves 10x latency reduction over autoregressive models while preserving ability to correct ASR mistranscriptions.
Visual-Aware Speech Recognition for Noisy Scenarios (2025.emnlp-main)

Copied to clipboard

Challenge: Existing audio-only models that use visual cues for transcription struggle in noisy environments.
Approach: They propose a method that correlates visual cues with noise sources to improve transcription by filtering speech from noise and predicting noise labels in video inputs.
Outcome: The proposed model improves transcription by correlating noise sources to visual cues in audio inputs.
DISCO: A Large Scale Human Annotated Corpus for Disfluency Correction in Indo-European Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing research on disfluency correction has primarily focused on English due to the unavailability of large-scale open-source datasets.
Approach: They propose to use an annotated human-annotated corpus to analyze disfluency correction in four important Indo-European languages to demonstrate the benefits.
Outcome: The proposed model improves BLEU scores by 5.65 points when used with a state-of-the-art machine translation system.
LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that pre-trained speech encoders and large language models can perform suboptimal performance on a range of spoken language processing tasks.
Approach: They propose to combine large-scale pre-trained speech encoders and large-language models for better performance on automatic speech recognition tasks.
Outcome: The proposed model can get an average of 49% WER reduction over the baseline model on 8 MLS testsets.
ATIR: Towards Audio-Text Interleaved Contextual Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Recent multimodal information retrieval research has focused on images, largely overlooking audio.
Approach: They propose an audio-text interleaved contextual retrieval task where queries can alternate between audio and text modalities.
Outcome: The proposed model significantly improves over baselines.
Fairness in Automatic Speech Recognition Isn’t a One-Size-Fits-All (2025.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained speech models like Whisper exhibit inconsistent group-level performance that varies across domains.
Approach: They fine-tune a Whisper model on the Fair-Speech corpus using basic fine- tuning, demographic rebalancing, gender-swapped data augmentation and a novel contrastive learning objective.
Outcome: The proposed method achieves stable, cross-domain fairness improvements without changes to the training data distribution and with minimal accuracy trade-offs.
SamróMur MilljóN: An ASR Corpus of One Million Verified Read Prompts in Icelandic (2024.lrec-main)

Copied to clipboard

Challenge: samrómur is a crowdsourcing web application designed to collect speech data for the advancement of language technologies in Icelandic.
Approach: They propose to use a crowdsourcing web application to collect and verify Icelandic speech data for automatic speech recognition (ASR) they introduce a dataset comprising one million audio clips from the application .
Outcome: The proposed system can produce high-quality speech data for Icelandic . the proposed system is based on a crowdsourced web application built on Mozilla's Common Voice .
Difference in Task Performance on Sparse Speech Representations (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for learning speech representations that are useful for a variety of downstream tasks have been extensively investigated in different domains.
Approach: They propose to train Autoencoders with varying sparsity levels using three SSL features and evaluate them on six tasks of SUPERB: speech enhancement, speaker identification, speech Emotion Recognition, phone recognition, automatic speech recognition and slot filling.
Outcome: The proposed model can be used to learn speech representations that are useful for a variety of downstream tasks.
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies show voice assistants do not perform equally well for everyone . however, research on demographic robustness of speech technologies is still scarce .
Approach: They propose a statistical method to detect demographic bias using a large dataset with controlled demographic tags.
Outcome: The proposed method shows statistically significant differences in performance across age, dialectal region and ethnicity.
Speech Recognition Corpus of the Khinalug Language for Documenting Endangered Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing tools to document endangered languages are limited due to data scarcity and the need for training.
Approach: They propose to use a speech corpus for Khinalug, an endangered language spoken in northern Azerbaijan, to create a model that can be used in language documentation scenarios.
Outcome: The proposed model achieves 6.65 CER points and 25.53 WER points in low-resource scenarios.
The Influence of Automatic Speech Recognition on Linguistic Features and Automatic Alzheimer’s Disease Detection from Spontaneous Speech (2024.lrec-main)

Copied to clipboard

Challenge: Existing biomarkers for AD diagnosis can only be applied to relatively small sample sizes due to limited availability, excessive costs and invasive nature.
Approach: They compare automatic speech recognition systems in terms of Word Error Rate (WER) using a publicly available benchmark dataset of speech recordings of AD patients and controls.
Outcome: The proposed method improves classification performance by replacing manual transcriptions with ASR output.
Unraveling Spontaneous Speech Dimensions for Cross-Corpus ASR System Evaluation for French (2024.lrec-main)

Copied to clipboard

Challenge: 'spontaneous speech' is a catch-all term used for situations like speaking with a friend, being interviewed on radio/TV or giving a lecture.
Approach: They propose to use four dimensions to describe spontaneous speech variation in automatic speech recognition systems.
Outcome: The proposed system can be used to predict the WER of speech recognition systems on face-to-face interactions.
MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) have advanced Chinese Classical Studies (CCS) but the audio dimension of CCS remains underexplored due to a lack of high-quality, domain-specific audio corpora.
Approach: They propose a 119-hour audio corpus comprising 22,000 audio samples to bridge this gap . it encompasses a diverse range of literary genres across six tasks .
Outcome: The proposed corpus encompasses a diverse range of literary genres across six tasks: Automatic Speech Recognition (ASR), Speech-to-Text Translation (S2TT), Speech Emotion Captioning (SEC), Spoken Question Answering ( SQA), Speech Understanding (SU), and Speech Reasoning (SR).
Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs (2026.acl-long)

Copied to clipboard

Challenge: Spontaneous speech is rarely fluent, and disfluencies can degrade readability and reliability . a sequence tagger first marks disfluent tokens, and these signals guide instruction fine-tuning .
Approach: They propose a multilingual correction pipeline where a sequence tagger first marks disfluent tokens . they add a contrastive learning objective that penalizes the reproduction of disfluency tokens.
Outcome: The proposed model improves readability and reliability of ASR transcripts in three languages . disfluencies can cause misinterpretations, incoherent responses, poor user experience .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations